Towards Efficient Searching on the Secondary Structure of Protein Sequences
نویسندگان
چکیده
Approximate searching on the primary structure (i.e., amino acid arrangement) of protein sequences is an essential part in predicting the functions and evolutionary histories of proteins. However, because proteins distant in an evolutionary history do not conserve amino acid residue arrangements, approximate searching on proteins’ secondary structure is quite important in finding out distant homology. In this paper, we propose an indexing scheme for efficient approximate searching on the secondary structure of protein sequences which can be easily implemented in RDBMS. Exploiting the concept of clustering and lookahead, the proposed indexing scheme processes three types of secondary structure queries (i.e., exact match, range match, and wildcard match) very quickly. To evaluate the performance of the proposed method, we conducted extensive experiments using a set of actual protein sequences. According to the experimental results, the proposed method was proved to be faster than the existing indexing methods up to 6.3 times in exact match, 3.3 times in range match, and 1.5 times in wildcard match, respectively.
منابع مشابه
Protein Secondary Structure Prediction: a Literature Review with Focus on Machine Learning Approaches
DNA sequence, containing all genetic traits is not a functional entity. Instead, it transfers to protein sequences by transcription and translation processes. This protein sequence takes on a 3D structure later, which is a functional unit and can manage biological interactions using the information encoded in DNA. Every life process one can figure is undertaken by proteins with specific functio...
متن کاملCSI: Clustered Segment Indexing For Efficient Approximate Searching On The Secondary Structure of Protein Sequences
Approximate searching on the primary structure (i.e., amino acid arrangement) of protein sequences is an essential part in predicting the functions and evolutionary histories of proteins. However, because proteins distant in an evolutionary history do not conserve amino acid residue arrangements, approximate searching on the proteins’ secondary structure is quite important in finding out distan...
متن کاملFTIR Investigation of Secondary Structure of Reteplase Inclusion Bodies Produced in Escherichia coli in Terms of Urea Concentration
Recent studies suggest that reducing the induction temperature would improve the quality of some recombinant inclusion bodies (IB) by providing a native-like secondary structure and leading to an improvement in protein recovery. This study focused on optimizing the solubilization condition of Reteplase, a recombinant protein with 9 disulfide bonds. The influence of lowering induction temperatur...
متن کاملFTIR Investigation of Secondary Structure of Reteplase Inclusion Bodies Produced in Escherichia coli in Terms of Urea Concentration
Recent studies suggest that reducing the induction temperature would improve the quality of some recombinant inclusion bodies (IB) by providing a native-like secondary structure and leading to an improvement in protein recovery. This study focused on optimizing the solubilization condition of Reteplase, a recombinant protein with 9 disulfide bonds. The influence of lowering induction temperatur...
متن کاملSearching on the Secondary Structure of Protein Sequences
Approximate searching on the primary structure (i.e., amino acid arrangement) of protein sequences is an essential part in predicting the functions and evolutionary histories of proteins. However, because proteins distant in an evolutionary history do not conserve amino acid residue arrangements, approximate searching on proteins’ secondary structure is quite important in finding out distant ho...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید
ثبت ناماگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید
ورودعنوان ژورنال:
- Fundam. Inform.
دوره 78 شماره
صفحات -
تاریخ انتشار 2007